Human Molecular Genetics
◐ Oxford University Press (OUP)
Preprints posted in the last 7 days, ranked by how well they match Human Molecular Genetics's content profile, based on 141 papers previously published here. The average preprint has a 0.12% match score for this journal, so anything above that is already an above-average fit.
Menon, R.; Khan, A. I.; Elangovan, D.; Kandadai, R. M.; Goyal, V.; Desai, S. D.; Joshi, D.; Kumar, H.; Wadia, P. M.; Mukherjee, A.; Kumar, N.; Mehta, S.; Geetha, T. S.; Sandeep, C.; Murugan, S.; Ayathu Venkat, M.; Shah, H. S.; Paramanandam, V.; Chandarana, M. v.; Yadav, R.; Dhamija, R. K.; Pal, P. K.; Biswas, A.; Gupta, R.; Borgohain, R.; Vedam, R. L.; Kukkle, P. L.
Show abstract
Parkinsons disease (PD) arises through disruption of multiple interconnected cellular processes, but the genetic contributions to these processes may differ across ancestries. We investigated functional convergence among genes harboring pathogenic or likely pathogenic (P/LP) variants and variants of uncertain significance (VUS) in a multicenter Indian cohort recruited through the Genetics of Parkinsons Disease in India Young Onset Parkinsons Disease project (GOPI YOPD). The cohort included 668 participants (463 males 69.3%) with a mean age at motor onset of 39.4+/-8.8 years. P/LP variants and VUS identified through previously reported whole-exome or whole genome sequencing were retained as separate evidential categories. The P/LP-associated gene set comprised 11 unique genes and the VUS associated set comprised 40 unique genes. Separate STRING functional-enrichment analyses evaluated Gene Ontology Biological Process, Molecular Function and Cellular Component terms, KEGG pathways, WikiPathways and STRING local network clusters. Terms meeting a Benjamini Hochberg false discovery rate threshold of <0.05 were organized into eight non-mutually-exclusive ontology/pathway categories. Gene to pathway mappings were subsequently projected to individual participants to estimate pathway representation and examine clinical associations. At least one reportable P/LP variant or VUS was identified in 336/668 participants (50.3%): 35 had a P/LP variant alone, 282 had VUS alone and 19 had a P/LP variant together with VUS in one or more additional genes. The most frequently represented categories were mitochondrial organization (247/336, 73.5%), autophagy related processes (228/336, 67.9%) and regulation of synaptic vesicle transport (201/336, 59.8%). PRKN was the most frequent P/LP-associated gene, occurring in 29/54 P/LP carriers, followed by PLA2G6 and PINK1. Lysosomal transport was represented exclusively by VUS-associated genes, particularly GBA1, VPS13C and LRRK2. Among P/LP carriers, additional VUS in distinct genes were not associated with age at onset (P = 0.81) or family history (52.6% versus 31.4%; P = 0.15). No pathway phenotype association remained significant after correction for multiple testing. Genetic findings in this Indian cohort converged across an interconnected mitochondrial autophagic lysosomal vesicular network, with different contributions from P/LP-associated and VUS associated gene sets. This study provides the first pathway resolved South Asian genetic profile and a framework for comparative studies across populations.
Saqib, M.; Chen, F.; Mistri, D. K.; Tan, L.; Wright, N.; Sarver, D. C.; Anders, R.; Aja, S.; Wong, G. W.
Show abstract
Trisomy 21 or Down syndrome (DS) affects multi-organ systems across the lifespan. The presence of an extra chromosome, along with genome dosage imbalance due to triplicated genes, contributes to the DS phenotypes. Of the DS mouse models, few are aneuploid with a freely segregating extra chromosome. We previously showed that the aneuploid Ts65Dn mice exhibit metabolic deficits consistent with the metabolic profile of DS. However, the genotype-phenotype relationships in Ts65Dn mice are complicated by the presence of triplicated genes unrelated to human chromosome 21 (Hsa21). To address this issue, we leveraged a refined model, Ts66Yah, where the extra triplicated genes in Ts65Dn have been removed. Deep phenotyping and multi-omics analyses showed that Ts66Yah mice develop pronounced and widespread metabolic disturbances. Despite sexual dimorphism in weight gain, body temperature, lipid and lipoprotein profiles, hepatic injury and adipose fibrosis, both male and female Ts66Yah mice share a common phenotype of pronounced glucose intolerance and insulin resistance, reduced mitochondrial respiratory capacity in visceral fat, altered serum inflammatory cytokine profile, and dysregulated serum and liver metabolomes. Pan-tissue transcriptomes also reveal signatures of immune activation, disrupted metabolic processes and cellular respiration, altered cytokine signaling, enhanced oxidative stress, and extracellular matrix remodeling. These combined changes across tissues disrupt metabolic homeostasis more severely in Ts66Yah than in Ts65Dn mice. Several phenotypes, including glucose intolerance, insulin resistance, tissue fibrosis, and oxidative stress were further exacerbated by an obesogenic diet. This foundational data establishes Ts66Yah as a valuable reference model for the mechanistic and comparative study of metabolic dysfunction in DS.
Chia, C.; Baker, K.
Show abstract
Obesity is a significant public health concern. Early-onset obesity in the context of rare disease can reflect genetically-mediated pathology or elevated susceptibility through indirect mechanisms. Mapping the diverse characteristics and needs of young people with obesity in the rare disease population is a first step toward mechanistic and translational research. We carried out a retrospective comparative analysis of demographic, genotypic, phenotypic and health service utilisation data for young people with obesity (cases: n=500) and without obesity (controls: n=11,444) from the UK 100,000 Genomes Project rare disease cohort. Cases and controls were recruited prior to genomic diagnosis, across clinical disorder categories. We observed significant association between socioeconomic deprivation and obesity risk. Young people with obesity had significantly higher utilisations of acute care and mental health services, indicating an overall higher health burden. A curated panel of 519 candidate obesity-associated genes demonstrated aggregate association with obesity, although no single gene reached significance. Phenotypic comparison between cases and controls highlighted increased multi-organ and neurological system involvement, highlighting the overlap between neurodevelopmental and obesity risks. Within the case group, we conducted cluster analysis to identify early-onset obesity groups with different phenotypic profiles, potentially arising from different causal pathways - this identified six obesity subgroups of interest, with differing involvement of neurodevelopmental and other systems. Our study confirms that obesity co-occurs with a wide range of factors within the rare disease population, and is associated with significant physical and mental health needs, requiring holistic lifelong care.
Montanez-Valverde, R. A.; Kim, V.; Duran-Luciano, P.; Yuan, Y.; Sofer, T.; Kaplan, R. C.; Gallo, L. C.; Talavera, G. A.; Perreira, K. M.; Daviglus, M. L.; Rosas, S. E.; Llabre, M. M.; Elfassy, T.; Li, X.; Isasi, C. R.; Rodriguez, C. J.
Show abstract
Background. The imprecision of current metrics to capture the complex genetic admixture and racial identity among Hispanic/Latino individuals in the United States [US] is a concern. We examined the relationship of self-reported race and genetic ancestry with hypertension [HTN] among Hispanics/Latinos. Methods. Cross-sectional study of the Hispanic Community Health Study/Study of Latinos (HCHS/SOL), including 10,586 Hispanic/Latino unrelated adults. Genetic ancestry: West African [AA], Amerindian [AI], and European [EA]. Self-reported race: White, Black, Native American, or Multiple/Missing (More than one race or Unknown/Not reported/Refused). HTN: systolic (SBP) [≥]130 mmHg, diastolic blood pressure (DBP) [≥]80 mmHg, and/or use of HTN medications. Age- and sex adjusted models were used. Results. Self-reported race was White (38{middle dot}6%), Black (3{middle dot}6%), Native American (4{middle dot}1%), and Multiple/Missing (53{middle dot}7%), with Unknown/Not reported/Refused representing 32{middle dot}7%. Black and White Hispanics/Latinos had the greatest AA (55{middle dot}7%) and EA (69{middle dot}3%) ancestries, respectively. Each 10% AA increase was associated with OR 1{middle dot}15, SBP beta +0{middle dot}9 mmHg, and DBP beta +0{middle dot}7 mmHg. Conversely, each 10% AI increase was associated with OR 0{middle dot}83, SBP beta -0{middle dot}4 mmHg, and DBP beta -0{middle dot}6 mmHg. HTN prevalence was highest among those with Black race or in the highest AA quantile (45{middle dot}6% and 48{middle dot}0%, respectively), and lowest among those with Native American race or in the highest AI quantile (37{middle dot}6% and 26{middle dot}7%, respectively). Conclusion. One-third of Hispanics/Latinos did not self-report race. Black or White self-reporting race did somewhat relate to AA or EA ancestry, respectively. HTN profiles were related to self-reported race and genetic ancestry in this admixed population.
Kwon, H. R.; Rackley, A.; Olson, L. E.
Show abstract
Autosomal dominant gain-of-function mutations in platelet-derived growth factor receptor beta (PDGFRb) cause overgrowth of the skeleton and other connective tissue in Kosaki overgrowth syndrome. However, the target cell type and signaling pathways underlying PDGFRb-driven overgrowth are unknown. Normal postnatal growth is controlled by pituitary-secreted growth hormone (GH), which activates the STAT5 transcriptional factor to upregulate insulin-like growth factor 1 (IGF1). To investigate the role of the GH-STAT5-IGF1 pathway in PDGFRb-related overgrowth, we generated mice with a PDGFRb gain-of-function mutation in skeletal and fibroblast lineages, which resulted in STAT5 activation and gigantism. Conditional deletion of Stat5ab in connective tissue lineages rescued skeletal overgrowth and keloid-like fibrosis in the skin. Conditional deletion of GH receptor (Ghr) did not rescue overgrowth, indicating the physiological activator of STAT5 is not required for overgrowth. However, deletion of Igf1, the STAT5 target gene, and its receptor, Igf1r, in connective tissue, rescued the overgrowth phenotype. These findings demonstrate a GHR-independent STAT5-IGF1 signaling pathway in mutant connective tissue cells, which mediates PDGFRb-driven overgrowth in mice and potentially in humans with similar PDGFRB mutations.
Efthymiou, S.; Tabata, K.; Dafsari, H. S.; Schober, E.; Latza, C.; Isaoglu, M.; Abuelrub, A.; Rad, A.; Firoozfar, Z.; Turchetti, V.; Lin, R. Q.; Maroofian, R.; Wiethoff, S.; Afzal, E.; Zafar, F.; Rana, N.; McRae, A. M.; Kaiyrzhanov, R.; Guliyeva, U.; Gulieva, S.; Melikishvili, G.; Lespinasse, J.; Vitobello, A.; Denomme-Pichon, A.-S.; Wentzensen, I. M.; Mefford, H. C.; Briere, L. C.; A Walker, M.; A High, F.; Sweetser, D. A.; Kendall, M.; Franchi, M.; Brown, M.; Latner, D.; Joset, P.; Ivanovski, I.; Alfadhel, M.; Alluhaydan, I.; Frederiksen, A. S.; Arriens, V.; Hanker, B.; Mankad, K.; Guerin, J
Show abstract
Pathogenic variants in RUBCN, encoding the Run domain Beclin-1 interacting and cysteine-rich domain-containing protein (Rubicon) have been implicated in autosomal recessive spinocerebellar ataxia 15 (SCAR15). However, the molecular mechanisms underlying disease pathogenesis remain poorly understood. Here, we report 18 individuals from 15 unrelated families harbouring biallelic RUBCN variants, who present with an aggressive neurodevelopmental disorder variably characterized by seizures, developmental delay, intellectual disability and movement abnormalities that cause regression, progressive brain atrophy and neurodegenerative features. Through functional characterization, we demonstrate that a subset of disease-associated putative truncating variants disrupt autophagy regulation. In Caenorhabditis elegans models, loss-of-function RUBCN variants result in an increased autophagic flux and impaired neuronal function, recapitulating key features in humans. Correspondingly, cellular assays reveal that nonsense and frameshift RUBCN variants lead to defective autophagy inhibition, underscoring a crucial role for RUBCN as a key negative autophagy regulator. Molecular dynamics simulations rank the eleven missense variants by structural effect, with p.Arg813Trp alone altering the target protein at both the local and the regional level and lying within the RAB7A-binding module that the truncating alleles remove altogether. Our findings establish and expand the RUBCN-related disorders as a clinically and molecularly distinct subset of autophagy-related diseases. By delineating both the genetic landscape and cellular consequences of Rubicon dysfunction, this study enhances our understanding of autophagy-related neurodevelopmental disorders and provides a foundation for future therapeutic investigations.
Satorres-Perez, E.; Castillo-Marco, N.; Igual, M.; Cordero, T.; Munoz-Blat, I.; Monfort-Ortiz, R.; Marcos-Puig, B.; Simon, C.; Garrido-Gomez, T.; Perales-Marin, A.
Show abstract
Background. In Europe, first-trimester combined screening with the Fetal Medicine Foundation (FMF) algorithm identifies women at increased risk of preeclampsia who may benefit from personalized aspirin prophylaxis. However, a substantial proportion of early-onset preeclampsia (EOPE) remains undetected at clinically acceptable specificity. Objective. To evaluate the first-trimester performance of MaiRa for early-onset preeclampsia (EOPE) risk stratification by benchmarking it against FMF screening in the same women, characterizing discordant patient-level classification profiles and exploring potential implementation strategies. Study Design. This secondary case-control analysis was nested within the prospective, multicentre PREMOM cohort [NCT04990141], which enrolled women with singleton pregnancies across 14 tertiary hospitals in Spain. First-trimester MaiRa and FMF risk estimates were evaluated in the same 126 pregnant women, comprising 99 uncomplicated controls and 27 EOPE cases, defined by disease onset before 34 weeks. Discrimination was compared using a stratified paired bootstrap analysis of the areas under the receiver-operating-characteristic curves. Performance was assessed at prespecified clinical thresholds, and detection rates were evaluated at fixed false-positive rates. Universal and contingent MaiRa implementation strategies were also evaluated. Results. MaiRa showed greater first-trimester discrimination for EOPE than FMF combined screening (AUC, 0.974 vs 0.900; P=.040) and consistently achieved higher detection rates across fixed false-positive rates. At false-positive rates of 5% and 10%, MaiRa detected 85.2% and 92.6% of EOPE cases, compared with 44.4% and 70.4% for FMF, respectively. Patient-level analysis demonstrated that MaiRa identified 12 of 27 EOPE cases (44.4%) classified as low risk by FMF; these pregnancies generally exhibited less abnormal conventional first-trimester profiles, including fewer maternal risk factors, lower mean arterial pressure and lower uterine artery pulsatility index, yet 8 of 12 (66.7%) subsequently developed severe EOPE. Exploratory implementation analyses showed that universal MaiRa screening achieved the highest EOPE detection, whereas a contingent strategy using FMF for triage and reflex MaiRa testing reduced molecular testing to 35.7% of pregnancies while maintaining 77.8% sensitivity and 97.0% specificity. Conclusion. MaiRa provided greater first-trimester discrimination for EOPE than conventional combined screening and detected additional pregnancies that later developed severe disease despite less abnormal conventional screening profiles. The findings suggest that maternal plasma cfRNA profiling captures biological alterations not fully reflected by combined first-trimester screening and support further prospective evaluation in an independent, unselected obstetric population. Key words: early-onset preeclampsia; first-trimester screening; cell-free RNA; liquid biopsy; Fetal Medicine Foundation algorithm; combined screening; risk stratification; aspirin prophylaxis.
Hasan, A.; Demidova, E. V.; Priyadarshini, P.; Czyzewicz, P.; Gathuka, L.; Murayama, T.; Zhou, Y.; Kiss, Z. A.; Shastry, R. K.; Andrake, M.; Hearne, G.; Devarajan, K.; Wu, C.; Shah, A.; Schultz, B. M.; Connolly, D. C.; Rosen, G. L.; Canadas, I.; Liu, J. C.; Burtness, B. A.; Smith, J. J.; Dunbrack, R. L.; Golemis, E. A.; Whetstine, J. R.; Meyer, J. E.; Arora, S.
Show abstract
Chemoradiotherapy (CRT) is the standard-of-care therapy for many solid malignancies, yet predictive biomarkers of treatment response remain limited. We identified a germline single nucleotide polymorphism (SNP) in an intrinsically disordered region of the lysine demethylase KDM3C/JMJD1C (p.S464T) that is associated with CRT outcomes in locally advanced rectal cancers (LARC) and head and neck squamous cell carcinoma (LA-HNSCC). In silico modeling with AlphaFold predicted S464T substitution influenced interaction between phosphorylated KDM3C and RNF8 FHA domain. In cellular models, conversion of S464 to T464 increased sensitivity to DNA-damaging agents. S464T substitution impaired damage-induced MDC1-RAP80 signaling and downstream RAP80-BRCA1 colocalization. SNP carrying cells impaired DNA repair causing genotoxic stress that is associated with increased cGAS-cGAMP innate immune signaling and increased apoptosis. Population analyses with the SNP highlighted an increase incidence of UV-induced skin and other cancers, linking inherited variation in the chromatin regulatory gene KDM3C to genome instability, cancer risk, and therapeutic vulnerability.
Mathews, R.; Bouyadjera, S. B.; Donegan, J. J.; Havird, J. C.
Show abstract
Mitochondria are central hubs for cellular metabolism and mitochondrial dysfunction is a hallmark of many chronic diseases. Consequently, changes in mitochondrial DNA copy number (mtDNA-CN), the number of mtDNA genomes per cell or tissue sample, are associated with diseases ranging from cancer and obesity to psoriasis and all-cause mortality. MtDNA-CN especially holds promise as a biomarker for neurodegenerative diseases, but whether and how mtDNA-CN changes with neurodegeneration is controversial. Here, we performed a systematic review and meta-analysis of 76 studies including 156 comparisons of mtDNA-CN in populations with or without a neurodegenerative disease to identify overall trends and potential moderators that explain variation among studies. Overall, mtDNA-CN was not statistically different with neurodegeneration, but heterogeneity among studies was extreme (I2 = 99.5%). The diagnosed disease explained the most variation. For example, Alzheimer's patients showed a 21% decrease in mtDNA-CN, but there was no change in mtDNA-CN with Parkinson's disease. Decreases in mtDNA-CN during neurodegeneration were also more extreme at older ages. Surprisingly, the tissue sampled for mtDNA-CN was not particularly influential, except for certain diseases. Studies published in earlier years also showed more extreme decreases in mtDNA-CN with neurodegeneration. Excessive heterogeneity persisted even after accounting for all moderators and their interactions (I2 = 85.7%). We conclude that the general perception of decreased mtDNA-CN with neurodegeneration is a vast oversimplification that may stem from legacy effects of early studies. However, mtDNA levels offer great promise as biomarkers for neurodegeneration, other diseases, and general health metrics, assuming appropriate complications can be considered.
Luo, X.; Syreeni, A.; Hill, C.; Smyth, L. J.; Dahlstrom, E. H.; Mutter, S.; Chen, Z.; Natarajan, R.; Pan, S.; Parton, A.; Jackson, H.; McKay, G.; Susztak, K.; Hirschhorn, J. N.; Florez, J. C.; Maxwell, A. P.; Groop, P.-H.; McKnight, A. J.; Sandholm, N.
Show abstract
Hyperglycaemia is a hallmark of diabetes and a major risk factor for diabetic kidney disease (DKD). However, the molecular consequences of long-term cumulative hyperglycaemia (CH) remain unclear. As a stable epigenetic modification, DNA methylation may capture past glycaemic exposure. Here, we assessed CH-associated DNA methylation in 1,245 participants with type 1 diabetes (T1D) from Finland and the United Kingdom-Republic of Ireland cohorts. We identified 17 CH-associated CpGs, with the strongest association at cg19693031 (TXNIP). Longitudinal analyses demonstrate that these CH-associated DNA methylation levels remain stable despite short-term glycaemic fluctuations, suggesting lasting epigenetic imprints of earlier metabolic control. Integrative analyses combining genomic, epigenetic, and proteomic data characterized these CpGs and potential target proteins. Mendelian randomization suggested a causal association between cg20853880 (KLF11) and DKD, supported by chromatin accessibility and kidney KLF11 expression. Our findings suggest that epigenetic changes contribute to metabolic memory and may mediate the effects of hyperglycaemia on DKD.
Clemsen, J. D.; Bockholt, H. J.; Adams, W. H.; Baker, B. T.; Bolton, J. L.; Calhoun, V. D.; Paulsen, J. S.
Show abstract
Background: The primary neuroanatomical site of Huntington-s disease (HD) pathology resides in the striatum and its atrophy identifies important disease progression from HD-ISS Stage 0 to Stage 1. Immune-associated proteins may capture variation in HD that is incompletely represented by markers of neuroaxonal injury. Objectives: To determine whether cerebrospinal-fluid myeloperoxidase contributes information about striatal volume loss beyond genetic disease burden and neurofilament light. Methods: Cross-sectional data from 88 persons with HD were analyzed. Cerebrospinal-fluid myeloperoxidase and neurofilament light were measured with a nucleic acid-linked immunosandwich assay. Normalized putamen volume was derived from structural magnetic resonance imaging. Linear regression adjusted for genetic disease burden and sex. Results: Higher neurofilament light was associated with smaller normalized putamen volume (standardized {beta} = -0.322, (P=.0066)). Higher myeloperoxidase was associated with larger normalized putamen volume after adjustment for genetic disease burden, sex, and neurofilament light (standardized {beta} = 0.183, (P=.0386)). Adding myeloperoxidase increased explained variance in striatal loss. Conclusions: Cerebrospinal fluid myeloperoxidase contributed modest incremental information about striatal volume in this cross-sectional sample. Independent longitudinal studies are needed to determine its biological source, temporal behavior, and potential biomarker value. Findings advance efforts to characterize multicomponent biological markers of HD.
Trindade Pons, V.; Gillespie, N.; Smit, R. A. J.; Arias, J. D.; Yin, X.; Berndt, S. I.; Oldehinkel, A. J.; van Loo, H.
Show abstract
Obesity is a growing public health challenge, with body mass index (BMI) influenced by both genetic and environmental factors. While the role of direct genetic transmission is well established, evidence for genetic nurture effects, in which parental genotypes impact offspring through the environment, has remained mixed. This study investigates direct genetic transmission and genetic nurture effects on BMI across ages, using parent-offspring trios and pairs from the Dutch Lifelines cohort study (N = 18,897 offspring, aged 8 to 67 years). We leveraged the latest multi-ancestry BMI polygenic score (PGS) to construct transmitted (PGS-T) and non-transmitted (PGS-NT) polygenic scores, where PGS-NT consists of parental alleles not passed on to offspring and serves as a proxy for genetic nurture. Linear mixed models showed a large effect of PGS-T on offspring BMI (Beta = 0.416, p < 0.001), corresponding to a 1.85 kg/m2 increase per SD increase in PGS-T. PGS-NT had a small but significant effect (Beta = 0.026, p = 0.013), consistent with a genetic nurture effect accounting for approximately 6.6% of the effect of direct transmission. Parent-of-origin analyses showed that maternal PGS-NT effects were larger than paternal effects. PGS-T interactions with age indicated that direct transmission effects increased in childhood and stabilized in adulthood, while PGS-NT effects remained stable across age. Our findings suggest that direct genetic transmission is the dominant influence on BMI, while results are consistent with small genetic nurture effects that are driven by the maternal side.
Jaholkowski, P.; Parker, N.; Sveen, I. O.; Wistrom, E. D.; Fominykh, V.; Szabo, A.; Parekh, P.; Frei, O.; Smeland, O. B.; O'Connell, K. S.; Djurovic, S.; Dale, A. M.; Shadrin, A. A.; Andreassen, O. A.
Show abstract
Recent large-scale studies have enabled new knowledge about genetic underpinnings of morphological and electrophysiological alterations of the retina. Variation in retinal traits, often of neurodevelopmental origin, have been linked to major psychiatric disorders (MPDs). Here, we investigate the genetic overlap between MPDs and key retinal traits to identify underlying molecular mechanisms. We obtained genome-wide associations studies data for bipolar disorder (BD), major depression (MD), schizophrenia (SCZ), and the retinal traits retinal nerve fibre layer thickness (RNFL), ganglion cell inner plexiform layer thickness (GCIPL), and vertical cup-disc ratio (VCDR). We estimated the number of trait-influencing variants shared between traits with MiXeR and identified shared genetic loci with condFDR. Subsequently, we examined the biological pathways of the genes mapped to shared loci. This revealed that GCIPL shared the most genetic variants with MPDs (~60%), followed by RNFL (~40%), and VCDR (~20%). The genetic variants shared between retinal traits and MPDs showed disorder-specific patterns with more pronounced overlaps of SCZ and BD with RNFL, and MD negatively correlated with GCIPL. Gene-pathway analysis highlighted the importance of GABAergic neurotransmission and a two-stage neurodevelopmental process in SCZ, whereas the role of mitochondria and a weaker developmental component were observed in BD. The results also implicated synaptic functioning and gene-expression processes in MD. Furthermore, polygenic analysis suggested that the genetic architecture of retinal traits can distinguish between MPDs. Our findings indicate shared genetic underpinnings between retinal traits and SCZ, BD, and MD, implicating altered neurodevelopment and neurotransmission underlying the retinal link to major psychiatric disorders.
Mansoor, R.; Minhas, A. S.; Thomas, A.; Mansoor, A. A.; McCambridge, A. H.; Dilts, C.; Eshak, J.; Govani, D.; Nylin, B.; Trinidad, J. C.; Kanaan, A. Y.; Kara, E.; Fielder, A.; Fielder, I.; Iglendza, A.; Mukatash, Y.; Pumnea, B.; Menzel, M. M.; Shabazz-Henry, A. L.; Niepielko, M. G.; Gao, M.
Show abstract
The QxxR motif is evolutionarily conserved within DEAD-box RNA helicases, including Drosophila Me31B and human DDX6, which post-transcriptionally regulate gene expression during animal development. A pathogenic H372R substitution (QxHR to QxRR) in the QxxR motif of human DDX6 has been associated with various developmental defects, but how this motif contributes to DDX6-family protein function remains unclear. Here, we used Drosophila Me31B as an in vivo model to investigate the QxxR motifs developmental role. We generated a Drosophila strain carrying the corresponding H333R missense mutation in Me31B and characterized its effects on female fertility, embryonic viability, germline development, and Me31B-associated molecular pathways. The me31BH333R mutation reduced female fertility in a gene dose-dependent manner, with homozygous mutant females being sterile. Embryos from the mutant females also exhibited primordial germ cell defects. Despite these developmental phenotypes, the me31BH333R mutation did not significantly alter Me31B protein abundance, global ovarian transcriptome or proteome profiles, or representative germ plasm mRNA and protein localization. In contrast, bait-normalized IP-MS analysis revealed altered enrichment of selected Me31B-associated proteins, including increased association of known Me31B interactors Trailer hitch (Tral) and Ypsilon Schachtel (Yps). These findings establish Me31BH333R as an in vivo model for investigating the conserved QxxR motif and suggest that disruption of this motif compromises development not through broad changes in gene expression, but potentially through altered composition or regulation of Me31B-containing ribonucleoprotein complexes.
Liu, H.; Mizani, M. A.; Zhao, Y.; Wood, A.; Inouye, M.; Price, A. L.; Jiang, X.; CVD-COVID-UK/COVID-IMPACT Consortium,
Show abstract
Predicting disease risk from prior diagnoses is fundamental to clinical decision-making, particularly during health emergencies such as the COVID-19 pandemic, when individuals with long-term conditions may be disproportionately vulnerable to adverse outcomes. Despite intense interest in developing models to predict disease risk from prior diagnoses (1-3), most prediction models do not estimate effects of each prior diagnosis on disease risk conditional on other diagnoses, limiting interpretability and clinical utility. We developed the Comorbidity Risk Score (CRS), trained on 13 million individuals (age 40-69) from linked electronic health record (EHR) datasets of the entire population of England, to predict COVID-19 hospitalisation and 87 other disease outcomes. CRS was trained at close to saturated sample size and precisely estimated the effects of 212 prior diagnoses on the 88 disease outcomes, conditional on all other prior diagnoses. Correlations of CRS effect sizes across outcomes (e.g. 0.76 for myocardial infarction vs. hyperlipidaemia) matched the corresponding genetic correlations (e.g. 0.79 for myocardial infarction vs. hyperlipidaemia), confirming that comorbidity architectures capture disease aetiology. On average, CRS identified 5% of the population with 3.4-fold higher disease risk, including myocardial infarction (4.4-fold), lung cancer (6.5-fold), and COVID-19 hospitalisation (6.3-fold). Using prior diagnoses alone, CRS outperformed state-of-the-art clinical COVID-19 models (4). Furthermore, CRS (N=13 million) substantially outperformed state-of-the-art AI (1) (N=0.5 million) and linear (3) (N=0.5 million) models in predicting disease risk, suggesting that training sample size outweighs model complexity. CRS attained near-perfect transferability across self-reported ethnicities (e.g., Black vs. White: AUROC ratio = 97.3%). Finally, CRS distinguished independently predictive comorbidities from indirect associations, e.g., lipid metabolism disorder was a strong predictor of myocardial infarction risk but not ischaemic stroke, after conditioning on other prior diagnoses. In conclusion, CRS provides a comprehensive resource for understanding the impact of comorbidities on COVID-19 and other future diseases, revealing disease aetiology while enabling powerful prediction of disease risk.
Yarmolinsky, J.; Cavallo, F. R.; Koskeridis, F.; Yu, X.; Bouras, E.; Richenberg, G.; Costantini, I.; Ray, D.; Woolf, B.; Karhunen, V.; Ellis, L.; Haycock, P. C.; Hemani, G.; Davey Smith, G.; Tsilidis, K. K.; Zuber, V.; McKay, J. D.; Dehghan, A.; Tzoulaki, I.
Show abstract
Confounding is a central challenge in observational studies. Here, we propose a framework for identifying confounders of two non-causally related traits by employing cross-trait pleiotropy analysis to detect genetic loci that affect both traits and multi-trait colocalisation to identify molecular phenotypes mediating these effects. We apply this approach to the analysis of C-reactive protein (CRP) - a non-specific marker of inflammation - and 10 inflammation-related cancers. In UK Biobank, higher pre-diagnostic CRP levels are associated with increased risk of multiple cancers, but bidirectional Mendelian randomization provides little evidence for a causal relationship. Cross-trait genetic analyses identify 92 loci with shared CRP-cancer effects including those with established roles in cancer and 50 novel loci such as RSPO3 (breast cancer) and GCKR (colorectal cancer). Integration with proteomic and single-cell transcriptomic data identified putative molecular mediators at 24 loci including plasma TLR1 levels in breast cancer and CD4+ T cell IRF5 expression in kidney cancer. Notably, 15 candidate effector genes encode targets of approved or investigational medications, including IL6, PDE4D, and CASP8, indicating potential opportunities for their repurposing for cancer prevention. The proposed approach provides a generalisable framework for leveraging non-causal phenotypic relationships to yield insights into disease mechanisms and therapeutic targets for disease prevention.
Zhu, J.; Baousi, A.; Morris, A. P.; Guo, H.
Show abstract
Standard polygenic risk scores (PRSs) are constructed based on additive genome-wide association study (GWAS) summary statistics. Nonlinear machine learning methods have been increasingly applied to construct PRSs directly from individual-level data, with the aim of improving predictive performance over standard PRSs through their ability to model non-additive genetic effects. However, their superiority across studies has been inconsistent, and the conditions under which they provide meaningful improvements remain unclear. We combined theoretical analysis, simulations and a real-world application to investigate when two widely used nonlinear machine learning methods, random forest and XGBoost, outperform standard PRSs. Theoretical analysis showed that standard PRSs can implicitly capture part of the genetic variance attributable to nonadditive genetic effects through their contributions to marginal SNP effects, thereby losing less information than commonly assumed. Although nonlinear models have a higher theoretical potential, their greater flexibility incurs a bias-variance trade-off that can limit predictive gains at finite sample sizes. Simulations showed that XGBoost outperformed the standard PRS only when the genetic architecture involves a sufficiently large proportion of interaction genetic variance concentrated across relatively few interaction effects and large training samples were available. Random forest consistently underperformed the standard PRS. In an application to ischemic heart disease prediction using UK Biobank data, XGBoost showed no meaningful improvement in predictive performance over the standard PRS, whereas random forest again performed worse. Together, these findings suggest that nonlinear machine learning do not uniformly outperform standard PRSs; rather, their relative performance depends jointly on genetic architecture and training sample size. Our study helps to reconcile the inconsistent results reported across previous studies and provides a framework for identifying settings in which more complex PRS models are likely to be beneficial.
pathak, s.; Richardson, T.; Sanderson, E.; Arora, N.; Strand, L.; Asvold, B. O.; Bhatta, L.; Brumpton, B.
Show abstract
Background: Higher Body Mass Index (BMI) is an established risk factor of sleep disturbance. It is not known if the effect is homogeneous across the lifecourse or if there is a particular time point in life that might be best to target. Methods: Two-sample Mendelian randomization (MR) was used to investigated the effect of childhood adiposity (adjusting on adulthood adiposity and obstructive sleep apnea (OSA)) on insomnia, morning chronotype, sleep duration, daytime sleepiness and daytime napping. Similarly, total, and direct effect of adulthood adiposity on these outcomes was explored. We used summary statistics from a genome-wide association study (GWAS) of UK Biobank for childhood and adulthood adiposity (n=453,169) and large-scale consortia of OSA (Million Veteran Program) (n=410,268), insomnia, and chronotype (23andMe) (n=1,978,022 and n=248,1000, respectively). Results: Two-sample univariable MR analysis provided no evidence of an effect of genetically predicted childhood adiposity on later life insomnia (Odds ratio (OR)= 0.94, 95% Confidence interval (CI)= 0.87, 1.03). Whereas, multivariable MR (adjusted for adulthood adiposity) analysis provide strong evidence of direct protective effect of genetically predicted childhood adiposity on later life insomnia (OR= 0.70, CI= 0.64, 0.77). Further, both in univariable and multivariable MR, a strong positive effect of increased childhood body size on morning chronotype was observed (OR= 1.16, CI= 1.01, 1.33 and OR= 1.36, CI= 1.15, 1.62, respectively) after accounting for adulthood body size. In both analysis the estimate did not change considerably after aditionally adjusting for OSA. However, childhood and adulthood adiposity found to be associated with OSA and OSA with insomnia. In both univariable and multivariable analysis, increased body size in adulthood increased the risk of having insomnia and a morning chronotype. Conclusions: The findings suggest that higher body size in childhood is not a risk factor for later life insomnia, whereas higher body size in adulthood was. Further, if healthy body size is maintained in adulthood, high childhood adiposity may decrease the risk of insomnia and increase the risk of being a morning person in later life. Keywords: childhood, adulthood, obesity, insomnia, morning chronotype, medelian randomization
Goroshchuk, O.; Koller, D.
Show abstract
Background: Endometriosis affects approximately 10% of reproductive-age women and is associated with substantial diagnostic delay and heterogeneous symptom presentation. Prior machine-learning prediction models have relied on comorbidity data alone or on small candidate-variant genetic scores, with inconsistent or incompletely reported performance. No study has combined a well-powered, multi-ancestry polygenic risk score (PRS) with environmental, reproductive, and symptom data in a single hybrid model. We developed and evaluated hybrid risk-prediction models integrating a genome-wide, multi-ancestry PRS with clinical and symptom data for endometriosis in the US-based All of Us Research Program. Methods: Among 69,376 participants (15,382 endometriosis cases, 53,994 controls) across six genetically inferred ancestry groups, we computed individual-level PRS values using PRS-CS weights derived from an independent, multi-ancestry GWAS. Five nested logistic regression, random forest, and XGBoost models progressively added age, ancestry, and within-ancestry genetic principal components (Model 1), environmental and reproductive factors (Model 2), symptom and comorbidity indicators (Model 3), all covariates combined (Model 4), and PRS x environment interactions (Model 5). Performance was assessed by AUROC in a held-out test set and 5-fold cross-validation, with class-weighted, Youden-optimized thresholds used for sensitivity, specificity, and predictive values; permutation importance identified top contributors. Pairwise AUROC differences were tested with a Holm-corrected DeLong-type test. Results: Discrimination improved from AUROC 0.63 (PRS, age, ancestry, principal components) to 0.72 for the full model, driven mainly by symptom and comorbidity data. XGBoost consistently outperformed logistic regression and random forest. The PRS ranked among the top individual predictors by permutation importance in nearly every model, alongside age, while genetic and demographic information alone gave only modest discrimination, and PRS x environment interactions did not improve on environmental factors alone. Threshold optimization yielded balanced sensitivity and specificity (~0.67/0.65) versus near-zero sensitivity at a default threshold. Conclusions: Combining the PRS with symptom and comorbidity data gave the best discrimination compared to solely a well-powered, multi-ancestry PRS as a predictor of endometriosis. This study clarifies both the promise and current limits of hybrid genetic-clinical prediction for endometriosis and points to symptom-based phenotyping, molecular subtyping, and external validation as priorities.
Belyea, M. M.; Shafiq, M.; Lass, J.; Much, C.; Liu, Z.; Kruse, N.; Haendler, K.; Sreenivasan, V.; Gelpi, E.; Siebels, B.; Ondruschka, B.; Spielmann, M.; Klein, C.; Trinh, J.; Glatzel, M.
Show abstract
Viral infections have long been proposed as environmental contributors to neurodegenerative diseases, including Parkinson's disease (PD), yet the molecular mechanisms linking infection and neurodegeneration are not well defined. Neuroinflammation and disruption of central nervous system (CNS) homeostasis have emerged as potential mediators. In this study, we used severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), the causative agent of COVID-19, as a model pathogen to investigate convergent molecular pathways between viral infection and PD. Single-nucleus RNA sequencing (snRNA-seq) was performed on post-mortem striatal tissue from 14 individuals stratified into four groups: COVID-19 only (COVID-19), PD only (PD), comorbid PD with COVID-19 (PD/COVID-19), and controls (Control). The PD/COVID-19 group exhibited an expanded astrocytic population and a pronounced interferon-associated molecular signature characterized by increased expression of canonical interferon-stimulated genes, including IFI44L (average log2FC= 3.9; adjusted p=2.3 x 10-373), IFI44 (average log2FC=2.9; adjusted p=8.0 x 10-266), ISG15 (average log2FC=3.1; adjusted p=1.2 x 10-197), and RSAD2 (average log2FC= 3.5; adjusted p=8.6 x 10-111). Pathway analyses demonstrated activation of innate immune and antiviral signaling pathways, particularly within microglia and astrocytes, including interferon signaling, pattern-recognition receptor pathways, and complement-associated responses. In parallel, genes involved in lipid metabolism, cholesterol homeostasis, synaptic maintenance, and neuronal signaling were reduced across disease groups. Proteomic analyses independently confirmed enrichment of antiviral and interferon-associated pathways and identified convergent suppression of sterol, cholesterol, and lipid metabolic processes. Our findings identify a convergent molecular signature linking PD and COVID-19, pronounced in comorbid individuals and characterized by interferon-driven innate immune activation, glial inflammatory responses, and dysregulation of lipid metabolic homeostasis. Collectively, the data support a model in which severe viral infection amplifies biological pathways already implicated in PD pathogenesis.